Back

Scientific Data

Springer Science and Business Media LLC

Preprints posted in the last 7 days, ranked by how well they match Scientific Data's content profile, based on 209 papers previously published here. The average preprint has a 0.15% match score for this journal, so anything above that is already an above-average fit.

1
Making Accelerating Medicines Partnership Data Findable and Interoperable through a Common Data Model: Extending OMOP for Multi-Source Multimodal Data

Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.

2026-09-02 genetic and genomic medicine 10.64898/2026.08.31.26361831 medRxiv
Top 0.2%
11.8%
Show abstract

SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.

2
Adaptive Post-Processing Recovers Most of the Gap to nnU-Net v2 in Head and Neck GTV Segmentation: A Paired Three-Arm HECKTOR 2025 Benchmark

Oyarzun Silva, R.; Hernandez Hernandez, P.

2026-08-31 radiology and imaging 10.64898/2026.08.28.26361649 medRxiv
Top 0.6%
4.2%
Show abstract

Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.

3
A single-session randomised crossover fNIRS study comparing three upper-limb mirror therapy task paradigms in healthy adults: a study protocol

Yang, T.; Wei, S.; Wang, Y.; Bai, D.

2026-09-02 rehabilitation medicine and physical therapy 10.64898/2026.08.28.26361691 medRxiv
Top 0.7%
3.5%
Show abstract

Background Mirror therapy (MT)-specifically paradigms using mirror visual feedback (MVF)-is widely used in neurorehabilitation; however, mechanistic implementations vary substantially in movement content, rhythmicity and attentional demands. This protocol describes an acute mechanistic, within-participant fNIRS screening study designed to compare three prespecified upper-limb mirror-therapy task paradigms and to quantify associated subjective experience after each condition in healthy adults during a single visit. Methods and analysis This is a single-centre, within-participant, randomised crossover study conducted at Wuhan Wuchang Hospital (Wuhan, China). Healthy adults aged 18-35 years will complete three task conditions once each in a counterbalanced order using a 3*3 Latin-square scheme: UMT1 (task-oriented rhythmic functional movement), UMT2 (open-ended free movement with auditory control), and UMT3 (non-functional rhythmic movement). fNIRS will be acquired using the NirSmart-6000A system during a standardised block design. The primary outcome is ROI-level HbO activation quantified as GLM-derived {beta} estimates within the prespecified primary ROIs (bilateral SM1/M1 and bilateral PMC). Secondary outcomes include ROI-level windowed {Delta}HbO (5-20 s post-onset relative to the immediately preceding rest; descriptive only), ROI-level {Delta}HbR, and post-condition subjective ratings (illusion, immersion, confusion and fatigue; 1-7 Likert). Condition effects will be analysed using linear mixed-effects models with fixed effects for condition and period and prespecified multiplicity-adjusted pairwise contrasts. Ethics and dissemination Ethics approval was obtained from the Ethics Committee of Wuchang Hospital Affiliated to Wuhan University of Science and Technology (Approval No.: 2025-112-01; approved on 2025-08-21). The study is expected to be minimal risk. Findings will be disseminated through publication of this protocol manuscript and subsequent results manuscripts and conference presentations. Trial registration number Chinese Clinical Trial Registry (ChiCTR2600116634). This study is conducted as a prespecified mechanistic sub-study under the overarching registered project.

4
Development of a new trauma dataset over 38 years from the Young Finns Study

Saarinen, A.; Asikainen, T.; Lehtimäki, T.; Raitakari, O.; Keltikangas-Järvinen, L.

2026-08-31 psychiatry and clinical psychology 10.64898/2026.08.26.26361417 medRxiv
Top 0.8%
3.2%
Show abstract

Background: Previous trauma research includes many limitations, such as the scarcity of pretraumatic health measurements and assessment of traumatic experiences with a broad scope across the lifespan. To respond to these gaps, we aimed to develop a new, prospective, population-based trauma dataset from childhood to middle age. Methods: We used the Young Finns Study that is a population-based, multi-generational, prospective study (n = 3596 for the main generation). It has started in 1980 (baseline assessment) and includes follow-ups in 1983, 1986, 1989, 1992, 1997, 2001, 2007, 2011/2012, and 2018-2020. From the 38-year follow-up and ten measurement points of the YFS, we collected all relevant trauma variables, including both free-format and structured questions that both the participants and their parents responded to. By a data-driven case-to-case analysis, we developed a scale to numerically capture variation in the quality of the experiences. Results: Our final dataset captured a total of 7769 traumatic experiences. We also developed the Traumatic Experience Severity Scale (TESS), including six subscales such as shamefulness, rarity, danger to life or health, effects on everyday life, human-made physical threat, and whether the target person was within or outside one's household. We also preprocessed the dataset to be later easily interleaved with other psychological, cardiovascular, and epigenetic variables of the YFS. Conclusions: We believe this new trauma dataset with thousands of experiences across the lifespan provides new opportunities to multidisciplinary, lifelong trauma research.

5
Reference-guided comparative genomics of seven Indonesian rice cultivars identifies conserved gene space and trait-associated sequence candidates

Purwestri, Y. A.; Wicaksono, A.; Nurbaiti, S.; Purba, N. T.; Retnaningati, D.; Restiani, R.; Kumalasari, N.; Nuringtyas, T. R.; Handayani, V. D. S.

2026-08-29 genomics 10.64898/2026.08.26.747264 medRxiv
Top 0.8%
3.1%
Show abstract

Indonesian rice cultivars represent valuable genetic resources, yet many remain poorly characterized at the genomic level. Here, we generated 95.40 Gb of PacBio HiFi sequence data from seven Indonesian rice cultivars and constructed cultivar-specific consensus genomes using the telomere-to-telomere Nipponbare reference AGIS1.0. Sequencing coverage ranged from 27.92x to 41.58x, and the resulting consensus genomes spanned 387.93-390.54 Mb, with BUSCO completeness of approximately 98.3-98.5%. OrthoFinder assigned 99.1% of predicted proteins to 40,737 orthogroups, including 27,514 core orthogroups represented across all seven cultivars, indicating a highly conserved predicted gene space within the reference-guided framework. Targeted analysis recovered 278 of 280 cultivar-by-locus combinations representing 40 genes or gene family entries associated with grain pigmentation, nitrogen and amino-acid metabolism, and starch properties. Comparative predicted protein analysis prioritized ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI as candidates for further investigation. Among 269 completed AGIS1.0-anchored promoter comparisons, 159 passed quality-control criteria, whereas 110 were flagged for gene-model, boundary, synteny, or structural concerns. Notably, these flagged comparisons accounted for more than 90% of the alignment-derived sequence variation, emphasizing the importance of rigorous quality control when interpreting apparent promoter divergence. Collectively, these reference-guided genomic resources provide a standardized framework for investigating sequence variation in Indonesian rice germplasm and prioritize testable coding and regulatory candidates for functional validation and future genomics-assisted crop improvement.

6
Deep Learning Frame Prediction for Abbreviated Low-Dose Dynamic PET Protocols on the PennPET Explorer

Courtens, J.; Muller, F. M.; Li, E. J.; Vanhove, C.; Vandenberghe, S.; Pantel, A. R.; Karp, J. S.; Daube-Witherspoon, M. E.

2026-08-31 radiology and imaging 10.64898/2026.08.25.26361357 medRxiv
Top 0.9%
2.5%
Show abstract

Dynamic positron emission tomography (PET) with long axial field-of-view (LAFOV) scanners enables multi-organ imaging and kinetic quantification beyond static (late-phase) imaging; however, the long times typically required for dynamic acquisitions remain clinically impractical. This study evaluates a deep learning (DL) framework to enable abbreviated dynamic PET acquisitions, comparing single-time-window (STW, early dynamic data only) and dual-time-window (DTW, early dynamic data plus a late 5-min static frame) protocols with early dynamic scan durations of 5-30 min and dose levels ranging from 360 MBq to 18 MBq. Seventeen 60-min dynamic [18F]FDG datasets were first motion-corrected using a staggered FALCON pipeline and then used to train and test a spatiotemporal DL model for autoregressive frame prediction. Performance was assessed across the full quantitative workflow, from DL-predicted frames and time-activity curves to organ-based kinetic modeling and voxel-wise parametric imaging in multiple tissues and two patient cohorts. DTW protocols consistently outperformed STW, better preserving late-phase kinetics. For a 15-min early dynamic scan, adding a late 5-min scan reduced mean absolute Ki difference from 23% (STW) to 17% (DTW) in the liver and from 26% to 15% in the thalamus. DTW + DL further reduced errors to [&le;]10% in the liver, thalamus, and breast lesion, and 16% in muscle. Our recommended protocol, 15-min early dynamic scan plus a 5-min late scan with DL, remained robust to up to a 5-fold dose reduction (~74 MBq). Overall, these findings support DL-enabled abbreviated, low-dose dynamic LAFOV PET as a clinically feasible approach for accurate kinetic quantification

7
The Stanford Knee Osteoarthritis PET/MRI Evaluation (SKOPE) Study Protocol

Goyal, A.; Vainberg, Y.; Shalit, R.; Gatti, A. A.; Kogan, F.

2026-08-31 radiology and imaging 10.64898/2026.08.26.26361112 medRxiv
Top 1%
2.4%
Show abstract

Purpose: The primary objective of the Stanford Knee Osteoarthritis PET/MRI Evaluation (SKOPE) study is to develop and evaluate a multimodal, dynamic [18F]NaF PET-MRI framework for characterizing whole-joint physiology and its relationship to osteoarthritis (OA) risk, pain, and disease progression. Specifically, we aim to integrate dynamic PET with quantitative and anatomical MRI, to characterize structural, compositional, and metabolic features across the knee and surrounding musculoskeletal system, evaluate acute tissue responses to exercise, and identify imaging biomarkers associated with OA risk, pain, and disease progression. Methods: The SKOPE study includes multimodal PET-MRI of the knee and surrounding musculoskeletal tissues, with imaging of the knee, tibia, ankle, thigh, hip, pelvis, and lumbosacral spine. Dynamic [18F]NaF PET is combined with conventional anatomical MRI and quantitative MRI techniques, including quantitative double-echo steady-state (qDESS) T2 mapping of cartilage, Dixon fat-fraction imaging, ultrashort echo time (UTE) T2* mapping of short-T2 tissues, UTE imaging of tibial bone, and zero echo time (ZTE) imaging for bone morphology and pseudo-CT generation. Additional MRI sequences characterize muscle composition, bone and joint anatomy, intervertebral discs, and regional vascular anatomy. Selected scans are acquired before and after a standardized exercise protocol to assess the acute physiological response of the joint. Automated segmentation is used to generate subject-specific masks of muscles, bones, vertebrae, and intervertebral discs. A subset of the MRI protocol is repeated at 1- and 2-year follow-up to assess longitudinal changes. Expected Impact: By combining dynamic bone metabolic imaging with quantitative measures of cartilage, menisci, muscle, bone, fat, vascular structures, and the spine and hip, the SKOPE protocol provides a whole-joint and multijoint framework for studying the structural, metabolic, and physiological processes associated with OA and pain. Exercise and longitudinal imaging further enable assessment of acute tissue responses and changes over time, supporting the development of quantitative imaging biomarkers for OA risk, pain, and disease progression.

8
SALRR: Scalable Analysis of Long-Read RNA-Seq Enables Comprehensive Transcriptome Profiling in Human Brain

Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.

2026-08-29 genomics 10.64898/2026.08.27.747499 medRxiv
Top 2%
1.5%
Show abstract

Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.

9
How Sex, Age, Adiposity, and Smoking Shape the Human Rib Cage: Evidence from 26,275 Whole-Body MRIs across the German National Cohort (NAKO)

Aicher, A.; Graf, R.; Kirschke, J.; Frauenfelder, T.; Ensle, F.; Menze, B.; Decker, J.; Kröncke, T.; Haubold, J.; Ringhof, S.; Bamberg, F.; Schmidt, C. O.; Wielpütz, M.; Leitzmann, M.; Willich, S. N.; Keil, T.; Niendorf, T.; Pischon, T.; Schlett, C.; Möller, H.

2026-09-03 radiology and imaging 10.64898/2026.09.01.26361964 medRxiv
Top 2%
1.5%
Show abstract

Rib-cage morphology is a determinant of thoracic biomechanics, ventilation, and injury response, yet statistical shape models (SSMs) of the rib cage have relied on small cohorts (~100s of individuals) imaged by clinical computed tomography, which over-represents injury and disease. We constructed a surface-based SSM of the complete 24-rib cage from 26,275 standardised whole-body magnetic resonance imaging (MRI) scans of adults aged 19-74 years from the population-based German National Cohort (NAKO). Ribs were segmented with a deep-learning pipeline (a rib-extended SPINEPS model), reconstructed as per-rib surface meshes, and brought into dense vertex-wise correspondence by Gaussian-process morphable registration in Scalismo; the aligned ensemble was summarised by generalised Procrustes analysis and principal component analysis (PCA). Fourteen per-rib geometric descriptors provided a quantitative cross-walk between the abstract PCA modes and named shape features, and associations with sex, age, body size and composition (including body-fat percentage), and smoking exposure were estimated by multivariable regression with Benjamini-Hochberg false-discovery-rate control. Shape variation was strongly concentrated: 28 modes captured 95% of the total variance, and the first three alone accounted for 69.4% (PC1, 42.6%; PC2, 16.3%; PC3, 10.5%) and admitted consistent anatomical readings - a sexually dimorphic axis (PC1), a slender-versus-stout body-habitus contrast (PC2), and a free-rib-size axis at ribs 11-12 (PC3). The sexes were nearly fully separated along PC1 (Cohen's d = 2.52). Body mass and body-fat percentage were the dominant modifiable correlates of rib-cage shape, whereas the association with cumulative smoking exposure was comparatively small. The model is released as a population-representative geometric reference for benchmarking and morphing donor-derived finite-element human-body models and for further large-cohort shape analysis.

10
Corpusome, a cross-body-site human microbiome corpus for representation learning

Xuan, H.; Huang, Y.; Bian, J.

2026-08-29 microbiology 10.64898/2026.08.28.747922 medRxiv
Top 2%
1.1%
Show abstract

Machine-learning models of the human microbiome are trained mostly on stool samples from single cohorts, limiting cross-body-site representation and cross-study generalization. Progress is constrained less by algorithms than by the absence of a harmonized multi-body-site corpus carrying the technical metadata needed to model, rather than ignore, batch structure. Here we release Corpusome, a harmonized two-tier cross-body-site human microbiome corpus for representation learning: a harmonized corpus of 187,546 human microbiome samples integrating standardized profiles from curatedMetagenomicData, the American Gut Project, and the EBI MGnify platform. Corpusome follows a two-tier design preserving both functional depth and cross-body-site breadth: a shotgun tier (22,588 samples, 93 studies) with species- and pathway-level profiles, and a 16S tier (164,958 samples, from a full pull of 708 MGnify studies) with genus-level profiles extending coverage to oral, skin, respiratory, and urogenital sites. It spans six body sites and two modalities, with harmonized metadata for batch-aware modelling. Body-site signal exceeds technical/source variance in the 16S tier by approximately 2.4-fold.

11
PyiTOL: reproducible Python workflows for iTOL annotation and taxonomic monophyly assessment

Zeng, Z.; Wang, Y.

2026-08-29 bioinformatics 10.64898/2026.08.27.747471 medRxiv
Top 2%
1.1%
Show abstract

Motivation: The Interactive Tree of Life (iTOL) is widely used to display and annotate phylogenetic trees, but managing its format-sensitive annotation files impede reproducible high-throughput analyses. Among the maintained Python packages and versions evaluated, none combined template generation, taxonomic monophyly assessment and iTOL batch operations. Results: PyiTOL validates inputs, generates 31 iTOL template schemas (22 accepted by the live batch uploader), performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay. On a topology-constructed benchmark, all calls matched prespecified labels for 4,389 groups; on a 700-genome tree, binary mono/non-mono calls agreed with ETE4 for 409 genera; 17,294 GTDB R232 genera were processed in about 17 s. Availability and Implementation: PyiTOL 1.0.3 (Python [&ge;]3.10; Linux, macOS and Windows) is MIT-licensed at https://github.com/ZengZichao/PyiTOL and archived with test data at Zenodo (https://doi.org/10.5281/zenodo.22106806).

12
What Matters Most: A Multi-Stakeholder Study of Outcome Domains in Lower-Limb Prosthesis Use

Ahmed, M. E.; Karlsson-Brown, S.; Koufaki, P.; Ahmadi, M.; Mico-Amigo, E. M.

2026-09-03 rehabilitation medicine and physical therapy 10.64898/2026.08.31.26361544 medRxiv
Top 2%
1.1%
Show abstract

Purpose: Lower-limb prosthesis use involves interacting physical, psychosocial, and device-related outcomes that may not be fully captured by conventional clinical assessment. This study aimed to develop and evaluate a stakeholder-informed framework of outcome domains relevant to meaningful everyday prosthesis use. Materials and Methods: A mixed-methods participatory design comprised a structured synthesis of selected clinically relevant content from five established patient-reported outcome measures; semi-structured interviews and importance and actionability ratings with 18 contributors (12 prosthesis users, four clinicians, and two industrial partners); and integration of the synthesis, qualitative, and rating findings. Interview records were analysed using reflexive thematic analysis, and ratings were analysed descriptively. Results: The resulting framework comprised four interrelated domains: Mobility, Physical Function, Psychosocial Wellbeing, and Prosthesis Experience. Mobility showed the clearest convergence across stakeholder perspectives. Prosthesis users showed the largest importance actionability gap for Prosthesis Experience (4.5 vs 3.0), whereas clinicians showed the largest gap for Psychosocial Wellbeing (5.0 vs 3.0). Interviews highlighted day-to-day variability in prosthesis use and the influence of confidence, fatigue, comfort, environmental conditions, social context, and device usability. Conclusions: Meaningful outcome assessment in prosthetic rehabilitation should extend beyond mobility alone to consider physical function, psychosocial wellbeing, and prosthesis experience within everyday contexts. The proposed framework provides a stakeholder-informed foundation for multidimensional outcome assessment in prosthetic rehabilitation.

13
REINA: A Recognize-Then-Infer Wearable-to-App AI Framework for Breast Cancer Rehabilitation

Zhuang, Q.; Mou, C.; Liu, B.; Fu, M. R.; King, G. W.

2026-08-31 rehabilitation medicine and physical therapy 10.64898/2026.08.29.26361725 medRxiv
Top 2%
1.1%
Show abstract

Breast cancer survivors frequently experience upper-limb impairments, making continuous monitoring essential for effective rehabilitation. We propose REINA (Recognize-Then-Infer Wearable-to-App AI Framework), a two-stage deep-learning approach for remote monitoring of motor function during breast cancer rehabilitation using wearable-device data. Inertial measurement unit (IMU) signals from wearable devices are first used to recognize physical activities via supervised learning, followed by an activity-specific recurrent neural network (RNN) to infer corresponding electromyography (EMG) signals. REINA establishes reliable inference of neuromuscular activity from wearable IMU data, enabling real-time, cost-effective assessment of motor function recovery in real-world settings.

14
Machine learning analysis of Autism phenotype data supports a four-dimensional continuum with three overlapping subtypes

Quigley, H.; Gardiner, B.; McDaid, L.; O'Donnell, C.

2026-08-31 psychiatry and clinical psychology 10.64898/2026.08.27.26361561 medRxiv
Top 2%
1.1%
Show abstract

Autism Spectrum Disorder (ASD) is a heterogeneous neurodevelopmental condition defined by differences in social communication and restricted, repetitive behaviours. As diagnostic criteria have broadened, ASD is now recognised across a wider range of individuals, raising key questions about its structure: does ASD have discrete sub-types, or is it better conceptualised as a continuous, possibly multidimensional, condition? We aim to explore whether a multidimensional continuum model more accurately captures the variability within ASD. We analysed a large SPARK phenotypic dataset of medical history and diagnostic surveys (background history, SCQ, RBS-R; n=36,710 individuals). We apply and compare two traditional statistical approaches, Factor Analysis and Gaussian Mixture Models, with a modern machine learning technique, the Variational Autoencoder (VAE). VAEs reconstructed unseen test data with ~4-fold better accuracy than Factor Analysis, and ~8-fold better accuracy than Gaussian Mixture Models. We identified four stable latent factors across 100 independently trained VAEs. These four dimensions provide an individual behavioural profile that can be visualized using radar-plots, offering a compact way to compare profiles at the person level. Through further analysis, we found evidence for 3 overlapping clusters or subtypes of ASD identified within the 4D latent space. This work aims to inform new ways of modelling ASD using a VAE that will be able to discern between a continuum or a clustered output and that go beyond binary diagnosis, instead reflecting the complex range of trait profiles, with implications for personalised diagnosis and intervention.

15
A Measurement-Based Care Strategy for Buprenorphine-Naloxone Treatment (Bup-MBC): Development of an EHR-Integrated Intervention

Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.

2026-09-01 addiction medicine 10.64898/2026.08.27.26361539 medRxiv
Top 3%
0.8%
Show abstract

Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.

16
An LLM enabled real-time estimation of seasonal influenza vaccine effectiveness from social media data

Pavia, M. J.; Amaro, I. F.; Xu, D.; Gonzalez-Hernandez, G.; Scotch, M.

2026-08-31 public and global health 10.64898/2026.08.28.26361670 medRxiv
Top 3%
0.8%
Show abstract

Influenza vaccine effectiveness (VE) is estimated from a limited number of clinics using a test-negative design. These standard estimates face geographic, temporal, and operational constraints. Using Twitter/X data, we applied few-shot chain-of-thought prompting to identify self-reported vaccination status and influenza test results, then implemented a test-negative-like design to estimate VE. Our estimates fell within the range of interim reports and could complement current systems, improving feasibility, timeliness, and scalability.

17
Prospective In-silico Simulation of the VESALIUS-CV Trial Using Biomedical Knowledge Graph and Real-World Data-Driven AI Modeling

Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361436 medRxiv
Top 3%
0.8%
Show abstract

Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.

18
Warming, thermal variability, and the 96% decline in childhood respiratory-infection mortality in China: a national time-series analysis of the Global Burden of Disease Study 2021 and the C-LSAT high-resolution climate dataset

Li, D.; Miao, Y.; Zhang, Y.; Chen, H.; Wang, X.; Shen, C.

2026-09-02 epidemiology 10.64898/2026.08.31.26361879 medRxiv
Top 3%
0.6%
Show abstract

Background Childhood respiratory mortality in China has fallen by over 90% in three decades alongside sustained national warming, yet national long-run evidence on temperature and child respiratory mortality is lacking. Methods We linked Global Burden of Disease (GBD) 2021 mortality estimates for China - lower respiratory infections (LRI), ages 0-19, and asthma, ages 0-24, 1990-2021 - with C-LSAT 0.5 deg gridded temperature data (1990-2019), aggregated nationally and to five climate zones. Four annual indicators (mean temperature, diurnal temperature range, seasonal amplitude, interannual variability) entered regressions of log mortality rates with Newey-West standard errors. A bootstrapped (500 resamples) quadratic model probed the minimum mortality temperature (MMT), with PM2.5-adjusted analyses and future-exposure, permutation, and detrended falsification tests. Results LRI deaths fell by 96.3% (330,194 in 1990 to 12,098 in 2021; 95% uncertainty interval 9,669-14,891) and asthma deaths by 94.9% (3,287 to 167), while mean temperature rose 0.364 deg C per decade and diurnal temperature range narrowed 0.092 deg C per decade. Baseline coefficients were large (mean temperature -1.696, SE 0.174; diurnal temperature range +2.408, SE 0.336; seasonal amplitude -0.162, SE 0.082; interannual variability +2.924, SE 1.514, per 1 deg C in log rate), but the future-exposure test failed and detrending nullified every coefficient: the associations are trend-level, and short-cycle causal effects are not identifiable. Nor was the national MMT identifiable - observed temperature support spans only 6.66-8.13 deg C, and the nominal turning point of 35.84 deg C is an extrapolation artifact (quadratic term p = 0.963). Within the observed range, warming and declining mortality moved in the same direction. Conclusions The 96% decline in childhood respiratory mortality cannot be attributed to warming. China sits on the low-temperature side of the optimum, and the marginal direction of future warming requires stronger designs to establish. The falsification framework offers a discipline for climate-health inference in China.

19
Best Practice Manufacturing and Quality Standards for Bacteriophage Therapy Products: Australian Consensus Statements

Watts, K.; Lin, R. C.; Lynch, S.; Warning, J.; Barr, J. J.; Ben Zakour, N.; Campbell, A.; Chan, J.; Collie, L.; Hedges, M.; Hudson, B.; Irwin, A.; Khatami, A.; Kicic, A.; Laucirica, D.; Lauter, C.; Ling, K.-m.; Ng, R.; Pavuk, N.; Rahmatullah, R.; Sinclair, H.; Tucker, E.; Vreugde, S.; Warner, M.; Velickovic, Z.; iredell, j.

2026-08-31 public and global health 10.64898/2026.08.26.26361487 medRxiv
Top 3%
0.6%
Show abstract

Objective As antimicrobial resistance (AMR) continues to threaten global public health, bacteriophage therapy products (BTPs) offer a promising alternative to conventional antimicrobials. However, translation into routine clinical practice requires best practice standards for manufacturing and quality control to ensure the consistent safety, quality, and reliability of personalised BTPs produced for individual patients or small cohorts. Design A modified Delphi methodology was used to develop consensus statements, engaging experts from Australia's National Bacteriophage Therapy Regulatory Working Group across the fields of clinical microbiology, phage biology, good manufacturing practice (GMP), regulatory science, and government. The process comprised three iterative phases: (1) structured statement development, (2) an anonymous REDCap survey, and (3) a hybrid consensus meeting. The strength of evidence and recommendations was assessed using the GRADE (Grading of Recommendations Assessment, Development and Evaluation) framework. Results Consensus was reached on 35 statements to provide best practice manufacture and quality control guidance for BTPs. These statements address requirements for phage identification and characterisation; define the point at which GMP-aligned processes commence for ubiquitous phages; outline quality control expectations for phage active pharmaceutical ingredient (pAPI) production and maintenance of BTP and host cell repositories. Additional guidance covers quality management systems, including documentation, traceability, and governance. Conclusion These consensus statements provide comprehensive best practice recommendations for the manufacture and quality control of BTPs in Australia. By promoting consistent, safe, and quality-assured approaches to personalised BTPs, they aim to facilitate clinical implementation while remaining aligned with existing international pharmacopoeial standards and regulatory frameworks.

20
ClinSeg: Robust Brain Segmentation for Clinically Acquired Pediatric MRI

Levitis, E.; Tregidgo, H. F. J.; Zimmerman, D.; Jung, B.; Karandikar, S.; Gardner, M.; Mattisson, P.; Kafadar, E.; Zapaishchykova, A.; Kann, B. H.; Sotardi, S. T.; Vossough, A.; Huang, H.; Billot, B.; Iglesias Gonzales, J. E.; Alexander, D. C.; Alexander-Bloch, A. F.; Seidlitz, J.

2026-09-02 pediatrics 10.64898/2026.08.28.26361643 medRxiv
Top 3%
0.6%
Show abstract

Clinical brain MRIs from pediatric health systems represent a viable resource for modeling early neurodevelopmental trajectories and studying neurodevelopmental risk in real-world populations. However, a limitation to date has been the performance of existing segmentation tools for measuring various brain phenotypes in clinical scans. In particular, many tools underperform in infant scans due to morphological and physical changes such as rapid myelination. Here, we introduce ClinSeg: a robust segmentation approach tailored to early-life clinical MRIs with variable orientation, resolution, and contrast. We leverage existing registration and synthetic data generation tools to construct a training corpus for a 3d U-Net spanning anatomical and contrast diversity, including scans with morphological abnormalities from a pediatric hospital. Validated against manual segmentations, ClinSeg outperforms existing models in infancy while matching them in childhood and adolescence. Finally, ClinSeg enables the construction of reference brain growth trajectories in 11,699 individuals from 0-21 years of age, leading to the detection of more nuanced age-related findings in clinical groups.